NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs

Feldman, A; Barak, L; Peng, J (May 2026, In Proceedings of the 15th Language and Resources Evaluation Conference 2026)

Full Text Available
The Multilingual Euphemism Benchmark: Datasets and Baselines for Pragmatic Language Understanding

Poh, W; Sammartino, J; Andrew, J; Kieraś, W; Zawadzka-Paluektau, N; Dilai, I; Barak, L; Peng, J; Feldman, A (May 2026, The Fifteenth Language Resources and Evaluation Conference (LREC) 2026)

Full Text Available
Euphemism Data - LREC 2026.

Poh, W; Sammartino, J; Andrew, J; Kieras, W; Zawadzka-Paluektau, N; Dilai, I; Barak, L; Peng, J; Feldman, A (February 2026, Euphemisms Benchmarks: LREC 2026)
Montclair NLP: Euphemisms

Feldman, A; Peng, J; Poh, W; Lee, P; Sammartino, J; Andrew, J; Trujillo, AC; Plancarte, CD; Ojo, EO; Shode, I; et al (March 2026, Euphemisms Project at Montclair State University)
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts

Xin, J; Yun, S; Peng, J; Choi, I; Ballard, J L; Chen, T; Long, Q (May 2025, https://doi.org/10.48550/arXiv.2505.19190)

Modality fusion is a cornerstone of multimodal learning, enabling information integration from diverse data sources. However, vanilla fusion methods are limited by (1) inability to account for heterogeneous interactions between modalities and (2) lack of interpretability in uncovering the multimodal interactions inherent in the data. To this end, we propose I2MoE (Interpretable Multimodal Interaction-aware Mixture of Experts), an end-to-end MoE framework designed to enhance modality fusion by explicitly modeling diverse multimodal interactions, as well as providing interpretation on a local and global level. First, I2MoE utilizes different interaction experts with weakly supervised interaction losses to learn multimodal interactions in a data-driven way. Second, I2MoE deploys a reweighting model that assigns importance scores for the output of each interaction expert, which offers sample-level and dataset-level interpretation. Extensive evaluation of medical and general multimodal datasets shows that I2MoE is flexible enough to be combined with different fusion techniques, consistently improves task performance, and provides interpretation across various real-world scenarios.
more » « less
Full Text Available
CAT s are Fuzzy PETs : A Corpus and Analysis of Potentially Euphemistic Terms

Gavidia, M.; Lee, P.; Feldman, A.; Peng, J. (January 2022, arXiv preprint arXiv:2205.02728.)

Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nevertheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, which may be useful for future applications. We also discuss the results of multiple analyses run on the corpus. Firstly, we find that sentiment analysis on the euphemistic texts supports that PETs generally decrease negative and offensive sentiment. Secondly, we observe cases of disagreement in an annotation task, where humans are asked to label PETs as euphemistic or not in a subset of our corpus text examples. We attribute the disagreement to a variety of potential reasons, including if the PET was a commonly accepted term (CAT).
more » « less
Full Text Available
Is self-supervised learning more robust than supervised learning?

Zhong, Y.; Tang, H.; Chen, J.; Peng, J.; Wang, Y.-X. (January 2022, Proc ICML Workshop on Pre-training)

Full Text Available
You Don’t Say....Linguistic Features in Sarcasm Detection

Ducret M.; Kruse L.; Martinez C.; Feldman A.; Peng, J. (January 2021, CLIC-IT 2021: Seventh Italian Conference on Computational Linguistics Bologna)

We explore linguistic features that contribute to sarcasm detection. The linguistic features that we investigate are a combination of text and word complexity, stylistic and psychological features. We experiment with sarcastic tweets with and without context. The results of our experiments indicate that contextual information is crucial for sarcasm prediction. One important observation is that sarcastic tweets are typically incongruent with their context in terms of sentiment or emotional load.
more » « less
Full Text Available
Distributed code for semantic relations predicts neural similarity during analogical reasoning

Chiang, J. N.; Peng, J.; Lu, H.; Holyoak, K. J.; Monti, M. M. (March 2021, Journal of cognitive neuroscience)
null (Ed.)
The ability to generate and process semantic relations is central to many aspects of human cognition. Theorists have long debated whether such relations are coarsely coded as links in a semantic network or finely coded as distributed patterns over some core set of abstract relations. The form and content of the conceptual and neural representations of semantic relations are yet to be empirically established. Using sequential presentation of verbal analogies, we compared neural activities in making analogy judgments with predictions derived from alternative computational models of relational dissimilarity to adjudicate among rival accounts of how semantic relations are coded and compared in the brain. We found that a frontoparietal network encodes the three relation types included in the design. A computational model based on semantic relations coded as distributed representations over a pool of abstract relations predicted neural activities for individual relations within the left superior parietal cortex and for second-order comparisons of relations within a broader left-lateralized network.
more » « less
Full Text Available
Pixel contrastive-consistent semi-supervised semantic segmentation

https://doi.org/10.1109/ICCV48922.2021.00718

Zhong, Y.; Yuan, B.; Wu, H.; Yuan, Z.; Peng, J.; Wang, Y.-X. (January 2021, International Conference on Computer Vision)

Full Text Available

« Prev Next »

Search for: All records